Back

Nature Methods

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match Nature Methods's content profile, based on 385 papers previously published here. The average preprint has a 0.41% match score for this journal, so anything above that is already an above-average fit.

1
Vipsania: Unsupervised Deep Gene Finding

Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.

2026-08-30 bioinformatics 10.64898/2026.08.26.747235 medRxiv
Top 0.1%
32.4%
Show abstract

Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.

2
Meso2EM: a cross-scale CLEM workflow linking mesoscale functional imaging to targeted electron microscopy

Oomoto, I.; Murate, M.; Sohn, J.; Tamura, M.; Hatada, S.; Egawa, N.; Odagawa, M.; Suga, M.; Kawaguchi, Y.; Murayama, M.; Kubota, Y.

2026-09-01 neuroscience 10.64898/2026.08.25.746890 medRxiv
Top 0.2%
30.6%
Show abstract

Meso2EM is a correlative light and electron microscopy workflow that transfers neurons selected from mesoscale functional images to targeted electron microscopy. We recorded Ca{superscript 2} signals from layer 2/3 neurons across a contiguous 3 x 3 mm cortical field in awake mice and reidentified a selected neuron after fixation and tangential sectioning. Lectin-labeled vascular architecture served as a shared landmark across in vivo two-photon imaging, confocal microscopy, laboratory micro-CT of resin-embedded tissue, and block-surface scanning electron microscopy, guiding focused-ion-beam scanning electron microscopy to the target cell body. The same progressive-targeting principle also supported serial ATUM-SEM reconstruction of an in vivo-tracked dendrite and serial transmission electron microscopy of optically selected dendrites from a patch-clamp-recorded Martinotti cell. Meso2EM therefore provides a practical route for preserving target identity across large changes in scale and specimen state while restricting electron-microscopy acquisition to a selected region.

3
JMod: Joint modeling of mass spectra for empowering multiplexed DIA proteomics

McDonnell, K.; Geiszler, D. J.; Wamsley, N.; Derks, J.; Sipe, S.; Cohen, Z. A.; Warinner, L. K.; Yeh, M.; Koo, E.; Leduc, A.; Zwang, T. J.; Specht, H.; Slavov, N.

2026-08-31 bioinformatics 10.1101/2025.05.22.655512 medRxiv
Top 0.4%
18.5%
Show abstract

Parallelization of data acquisition substantially increases the throughput of mass spectrometry-based proteomics. However, parallelization also increases the density of mass spectra and consequently the overlap between ions, frustrating their analysis. To improve sequence identification and quantification from such spectra, we developed an open-source software for Joint Modeling of mass spectra (JMod). JMod models overlapping peaks as linear superpositions of their components in both MS1 and MS2 space, which permits multiplexed DIA with smaller mass offsets to increase the multiplexing capacity and thus proteomics throughput for a given plexDIA tag. This enables 9-plexDIA using 2 Da offset PSMtags, increasing throughput 9-fold while preserving quantitative accuracy and coverage depth. Furthermore, we use JMod to deconvolve simultaneous labeling by mass tags and heavy amino acids, thus increasing the throughput of metabolic pulse experiments measuring protein synthesis and degradation rates in single cells from mouse liver. By supporting enhanced decoding of highly multiplexed DIA spectra, JMod provides an open and flexible software that increases the throughput of sensitive proteomics.

4
RegimeFormer: A Large Protein Model of Global Perturbation Regimes

Ma, S.; Chai, Y.; Wu, Y.; Zhang, Q.; Yuan, Y.; Zhao, K.; Chen, Z.; Wang, H.; Cao, S.; Yu, X.; Han, X.; Liu, Y.; Liu, Y.; Zhu, T.; Tao, D.

2026-08-30 bioinformatics 10.64898/2026.08.26.747182 medRxiv
Top 0.5%
18.2%
Show abstract

Protein language models organize sequence and structure at scale, but a global representation of how proteins respond to mutation remains lacking. We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, constructed by harmonizing and indexing 202,556,313 non-redundant protein sequences across the tree of life. A diversity-preserving one-million-protein subset provides the high-resolution training and inference layer, with 995,995 proteins yielding residue-level summaries across 407,048,356 residues and substitution-specific predictions available on demand. Across experimental deep mutational scanning, molecular benchmarks, structural confidence and evolutionary constraint, RegimeFormer identifies reproducible protein-level perturbation regimes that organize residue fragility, adaptability and predictive uncertainty. Regime conditioning improves substitution-specific prediction, with the largest relative gains under unseen-protein, unseen-family and low-homology evaluation. RegimeFormer-derived molecular priors further improve downstream transcriptomic and drug-response modelling. Together, RegimeFormer and RegimeAtlas provide a scalable framework for mapping, predicting and querying protein perturbation landscapes across global sequence space.

5
Calibration-free compression brings Evo 2 to its full million-token context on a single GPU

Patsakis, M.; Tzanakakis, A.; Georgakopoulos-Soares, I.

2026-09-01 bioinformatics 10.64898/2026.08.28.747902 medRxiv
Top 0.5%
17.2%
Show abstract

Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.

6
Audited vibe coding suggests partial fetal-like convergence of tumor proteomes

Meyer, J. G.

2026-08-31 cancer biology 10.64898/2026.08.26.745609 medRxiv
Top 0.6%
15.0%
Show abstract

The balance between how much human tumors recapitulate fetal tissue programs versus lose adult tissue identity remains unresolved. I used audited vibe coding, a human-mediated, cross-model critique-and-refinement workflow, to re-analyze a public pan-cancer proteomic atlas. A primary large language model wrote and executed the analysis under scientific direction, while a separate model family audited the code, outputs and claims; findings were returned for correction across seven versioned releases. Among 229 tumor-adjacent pairs in seven organs, tumor-minus-adjacent proteomic change partially aligned with reverse fetal-to-adult maturation (organ-balanced cosine, 0.240; 95% interval, 0.138 to 0.335), with positive alignment in 189 of 229 patients (82.5%). The organ-balanced projection coefficient was 0.195 (95% interval, 0.069 to 0.244), indicating movement along only part of the developmental distance. Although reverse maturation overlapped adult-identity loss, a positive developmental component remained after identity loss entered first (0.203; 95% interval, 0.129 to 0.239). Suppression of adult-high proteins contributed to more positive alignment than reactivation of fetal-high proteins. The vibe coding audits identified substantive defects. A common-mask correction reduced the matched-organ advantage from 0.074 to 0.059; a missing-value correction barely changed aggregate geometry but replaced 5 of the top 40 liver contributors; and coupled resampling repaired uncertainty accounting without changing patient scores. As with any single report, the "vibe reanalysis" biological results are candidate discoveries pending independent replication. The workflow is a single feasibility case, not a reliability benchmark, and shows how conversationally generated analysis can be made more inspectable when model-written code is treated as untrusted, versioned and subject to separate-model critique and executable checks.

7
MechanoMaST - a multimodal pipeline for spatially registering mechanical and transcriptomic tissue data

Decker, L.; Olisov, D.; Schleussner, N.; Wiethoff, H.; Schmidt, T.; Nienhueser, H.; Pausch, T. M.; Korbel, J. O.; Diz-Munoz, A.

2026-08-31 biophysics 10.64898/2026.08.29.747727 medRxiv
Top 0.6%
15.0%
Show abstract

Spatial-omics workflows enable molecular analysis within tissue spatial context. Despite the prognostic value of tissue stiffness, these approaches have not incorporated direct, mechanical measurements. This omission reflects several challenges, including sample requirements, low throughput, specialized equipment, and complex data registration. Here, we introduce mechanoMaST (mechanics mapped to spatial transcriptomics), the first workflow to combine absolute mechanical measurements with spatial-omics. It pairs atomic force microscopy-based nanoindentation stiffness maps with spatial transcriptomics maps from adjacent tissue cryosections. The two modalities are then computationally co-registered to enable direct spatial correlation at 100 um resolution, with mapping accuracy quantified through error propagation, providing ground-truth mechanical data directly linked to spatial gene expression. We demonstrate mechanoMaST in human colorectal cancer liver metastasis, generating a spatial resource from 10 patients and revealing a four-gene stiffness signature. mechanoMaST is readily adaptable to other tissues across development and disease, and extendable to additional spatial-omics modalities in adjacent sections.

8
Data-driven spectroscopic dictionaries and detector-calibrated inference for photon-limited Raman hyperspectral imaging of living cells

Yagi, S.; Sagami, N.; Eshima, I.; Hiramatsu, K.

2026-09-01 cell biology 10.64898/2026.08.31.748229 medRxiv
Top 0.6%
14.7%
Show abstract

Label-free Raman imaging of living cells is photon limited: at exposures compatible with cellular dynamics, single-pixel spectra carry about one count per channel on a dominant smooth background. We present an unmixing framework in which the decoder of a physics-constrained autoencoder is restricted to a data-driven spectroscopic dictionary: band centers,widths, and pseudo-Voigt shapes are measured from the dataset and fixed, and the network learns only nonnegative band amplitudes, a smooth B-spline background, and a per-pixel gain.First, on slit-scanning images of HeLa cells (532 nm) the dictionary yields spike-free component spectra that read as band tables, including a resonance-enhanced cytochrome-c-associated component matching literature spectra, and the most stable decomposition against the component number. Second, the dictionary and initialization calibrated at 1 s exposure perline transfer to 100 ms per line (12 s sweeps): cytochrome-c spectral identity survives a single sweep (correlation 0.92) while its map remains photon limited; the dictionary provides spectral physicality, and the transferred initialization prevents a structural collapse that global map correlations miss; in a measurement-derived phantom the dictionary estimator holds thecytochrome-c spectrum to 17-19{degrees} spectral angle at 100 ms, where classical factorizations and free decoders lose it (55-64{degrees}). Estimation on the count-equivalent detector output uses a calibrated shifted-Poisson quasi-likelihood. Third, evaluation must be time matched:correlation against a separately acquired reference saturates through slow specimen drift and acquisition mismatch rather than photon noise, and the self-consistency of learned denoisers is inflated by shared bias; time-matched self-consistency and independent cross-checks areproposed.

9
RECON infers regions of interest from H&E images and reconstructs whole-slide molecular profiles at single-cell resolution

Yang, X.; Hao, N.; Zhao, R.; Angel, S.; Tan, Y.; Lian, C. G.; Zhou, L.; Olson, D.; Yu, K.-H.; Ruiz de Luzuriaga, A.; Wan, G.

2026-09-01 bioinformatics 10.64898/2026.08.25.747122 medRxiv
Top 0.8%
11.9%
Show abstract

Spatial omics technologies resolve molecular expression and spatial architecture at single-cell resolution, but profiling whole slides remains costly. In practice, only a few regions of interest (ROIs) are profiled, leaving the rest of the tissue unmeasured. S2-omics was the first framework to unify ROI selection with out-of-ROI prediction, but it operates on superpixels rather than individual cells and predicts discrete cell types rather than continuous molecular profiles. Superpixel-based representations do not explicitly preserve cell boundaries, while categorical cell-type labels cannot quantify molecular expression within cells. Here we present RECON, a two-stage framework that performs ROI inference and whole-slide molecular reconstruction at single-cell resolution, predicting both continuous molecular profiles and discrete cell-type labels. In the first stage, RECON extracts morphological and microenvironmental features from individual cells to identify a representative ROI for spatially resolved single-cell molecular profiling. In the second stage, RECON trains deep learning models on molecular measurements acquired within the selected ROI and reconstructs transcriptomic or proteomic profiles for all remaining cells on the slide. Benchmarked against pathologist annotations, RECONs ROI selection outperforms the superpixel-based S2-omics approaches (IoU: 0.75 versus 0.64). For transcriptomics, refining the modeling unit from superpixels to single cells improves per-gene Pearson correlation by 22%. For proteomics, RECON surpasses the current state-of-the-art method, ROSIE, across all 16 markers, with a median per-cell Pearson correlation of 0.91 versus 0.84. Moreover, RECON delineates tumour boundaries and regions with distinct immune-cell densities, and highlights candidate tertiary lymphoid structures. Together, these results demonstrate that RECON enables informative ROI selection and whole-slide molecular reconstruction at single-cell resolution for both spatial transcriptomics and spatial proteomics.

10
Pre-FIB Layer-Mapping Cryo Tomography (PLCT) for Depth-Resolved in Situ Structural Analysis of Multilayered Tissues

Wang, F.; Lin, X.; Rao, B.; Lai, X.; Yu, L.; Sun, F.; Qu, J.; Zhang, J.

2026-08-30 neuroscience 10.64898/2026.08.25.746966 medRxiv
Top 0.8%
11.8%
Show abstract

Cryo-electron tomography (cryo-ET) enables near-native visualization of subcellular architectures, yet applying it to moderately thick, multilayered tissues such as the retina is hampered by inadequate vitrification and inaccurate depth-targeting. Here, we developed PLCT, an integrated approach combining modified high-pressure freezing, cryo-ultramicrotome trimming, and plasma-based cryo-FIB milling to overcome these barriers. PLCT reliably vitrified <100 m retinal strips with minimal ice artifacts, navigates precisely to the outer plexiform layer using morphological landmarks, and produces high-quality lamellae suitable for high-resolution cryo-ET. Subtomogram averaging (STA) analysis identified microtubules at 16.33 [A] within retinal horizontal cell processes. Importantly, STA also resolved a 10-nm-diameter filamentous structure at 24.81 [A] in the same processes, featuring six peripheral strands surrounding an elongated central density with continuous intervening cavities, an architecture consistent with intermediate filaments. Together with its native localization and immunoreactivity, these features collectively identify the filaments as neurofilaments. Separately, 3D reconstruction of synaptic ribbons uncovered a previously unrecognized "mahjong tile"-like fine ultrastructure. These results demonstrate that PLCT-produced lamellae are of sufficient quality to support structural analysis in native tissue. Although demonstrated on retinal photoreceptor synapses as a proof-of-principle, PLCT is inherently generalizable, with its depth-navigation and vitrification strategies directly applicable to any multilayered tissues. This work establishes PLCT as a robust, reproducible platform for depth-resolved in situ cryo-ET of multilayered tissues.

11
Genome-scale label-free imaging reveals cellular physiology encoded in bacterial collective architecture

Mellick, S. N. S.; Derringer, J. J.; Boyes, D.; Croteau, G.; Burke, M.; Gifford, S.; Stark, D. J.; Mike, L. A.; Turecki, S.; Carja, O.; Mikheyeva-Bridges, I. V.; Bridges, D. A.

2026-08-31 microbiology 10.64898/2026.08.30.748126 medRxiv
Top 1%
9.8%
Show abstract

DNA sequencing unified microbial genotyping into a single, comprehensive readout, yet phenotyping remains a slow and fragmented endeavor. Here, we introduce Microbial Phenotyping Using Low-magnification Label-free Imaging (PULLI), a computer vision platform that extracts microcolony and population-level phenotypes from brightfield timelapses of liquid culture growth. Using PULLI, we screened a genome-scale Vibrio cholerae mutant library, recording more than 200,000 images, which revealed that core bacterial pathways shape community architecture. Functionally related mutants converge in appearance, allowing us to resolve processes as distinct as biofilm formation, motility, central metabolism, cofactor biosynthesis, and envelope composition using a single approach. We further show PULLI can be used to determine a drug target, characterize other pathogens, and classify bacterial species. Our results show that bacterial multicellular development is an interpretable signature of genotype-phenotype relationships, which can be captured from simple brightfield timelapses. We release the PULLI pipeline and an interactive atlas of community forms.

12
ChemIntelligence Enables Antibody-Free, Ultra-Low-Input Profiling of Lysine Lactylation and Diverse Acyl-Proteomes

Shao, C.; He, Z.; Yuan, Q.; Giurcoiu, V.-G.; He, X.; Cao, X.; Huang, H.; Zhang, Y.; Zhang, Y.; Wang, D.; Jiang, Q.; Guo, Z.; Hao, H.; Wilhelm, M.; Ye, H.

2026-08-31 biochemistry 10.64898/2026.08.28.746934 medRxiv
Top 1%
9.8%
Show abstract

Lysine acylations, including lactylation (Klac), are pivotal regulators of cellular physiology. However, their analysis is currently bottlenecked by antibody enrichment strategies that suffer from sequence bias and require milligram-scale protein inputs, severely precluding the profiling of scarce clinical biopsies and rare cell populations. Here we present ChemIntelligence, an acyl-NHS chemistry-empowered derivatization strategy that rapidly generates unprecedented acylation-specific spectral libraries, exemplified by over 2.5x10^9 human Klac peptides, enabling cross-species reference atlases. Integrated with Prosit-based rescoring, these libraries substantially increase Klac identifications across diverse proteomic datasets. Leveraging this spectral resource, we devised ChemIntelligence Scope, a reproducible, multiplexed parallel reaction monitoring (PRM) platform that quantifies hundreds of Klac peptides per injection from as little as ~200 ng of cell lysates, clinical biopsies, and even true single cells - revealing functional Klac signatures inaccessible to conventional methods. The ChemIntelligence pipeline also extends seamlessly to lysine nicotinylation, underscoring its broad adaptability for discovering and profiling new acylations. Together, these chemical and computational advances establish a scalable, antibody-free framework for acyl-proteome mapping that overcomes input constraints and enables deep functional insights from otherwise intractable biological samples.

13
scPyviewer: a Python-native interactive viewer from AnnData single-cell data

Xuan, H.; Huang, Y.; Bian, J.; Liu, X.

2026-08-31 bioinformatics 10.64898/2026.08.26.747418 medRxiv
Top 1%
9.7%
Show abstract

Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.

14
Spatial Transcriptomics As Rasterized Image Tensors (STARIT) characterizes cell states with subcellular molecular heterogeneity

Velazquez, D.; Hallinan, C.; An, R.; Clifton, K.; Fan, J.

2026-09-01 bioinformatics 10.64898/2025.12.18.695193 medRxiv
Top 1%
7.8%
Show abstract

Abstract Imaging-based spatially resolved transcriptomics (imSRT) technologies provide high-throughput molecular-resolution spatial characterization of genes within cells. Conventional analysis methods to identify cell-types and states in imSRT data rely on gene count matrices derived from tallying the number of mRNA molecules detected for each gene per segmented cell, thereby overlooking subcellular heterogeneity that can be useful in defining cell states. To take advantage of the molecular-resolution information in imSRT data and potentially identify cell-states based on subcellular heterogeneity, we developed STARIT (Spatial Transcriptomics As Rasterized Image Tensors). STARIT converts transcripts within segmented cells in imSRT data into an image-based tensor representation that can be combined with deep learning computer vision models for downstream analysis. Using simulated and real imSRT data, we demonstrate that STARIT distinguishes transcriptionally distinct cell-types and further separates cell states based on subcellular transcript localization, which conventional gene count analysis fails to capture. By providing a standardized framework to encode subcellular molecular information in imSRT data, STARIT will enable deeper insights into subcellular heterogeneity and enhance the identification and characterization of cell-types and states that are overlooked by gene count representations.

15
Democratizing three-dimensional surface phenotyping: an open structured-light platform reveals and removes the projection bias in biological imaging

Gentsch, G. J.; Guo, M.; Platz, A.; Brehm, G.; Hennings, J. C.; Huebner, C. A.; Stark, A. W.; Franke, C.

2026-08-31 bioengineering 10.64898/2026.08.30.748077 medRxiv
Top 2%
6.8%
Show abstract

Surface phenotyping underpins plant science, preclinical animal research and entomology, yet across all three the measurement is almost always a photograph, which records a projection and not the surface itself. Here we present the Gentschinator3000, an open structured-light platform that brings high-end metric surface measurement within reach of laboratories with no optics expertise, combining documented open hardware, open reconstruction software and analysis workflows for under 4000 Euro in components. It resolves a planar reference to 45 m local flatness, registers full rotations to a loop closure of 156 m, and performs stably across acquisition ranges that we define. Applying one workflow to a leaf before and after desiccation, to murine anatomy and to a spread lepidopteran, we find that projection underestimates surface area by 11 to 41 %. That error grows with the condition under study, with the evaluation scale and with the direction of view, so it can confound phenotype comparisons dramatically. In murine limbs a 15-degree change of viewing direction shifts a projected inter-segment angle by up to 23.2 degrees, while the three-dimensional angle does not move. Projection geometry can therefore contribute as much to a measured phenotype as the biology it is meant to quantify.

16
OMICON: a community resource for studying gene coexpression networks in normal and neoplastic human brain samples

Eliscu, R.; Kang, G.; Schupp, P. G.; Brody, D. J.; Hariharan, N.; Shamsian, S.; Oldham, M. C.

2026-09-01 neuroscience 10.64898/2026.08.25.747141 medRxiv
Top 2%
6.5%
Show abstract

Genome-wide coexpression analysis of intact tissue samples is a powerful approach for identifying reproducible signatures of cell types and states, since it can survey vast numbers of individuals, cells, and transcripts. However, it can be difficult to optimize gene coexpression network construction and compare results from independent analyses. To address these challenges, we developed OMICON (theomicon.ucsf.edu) for research on human brain gene coexpression networks. OMICON contains gene expression data from >17K normal and neoplastic human brain samples with standardized metadata. Systematic analysis of independent datasets identified >250K gene coexpression modules, which were characterized and compared via enrichment analysis with >40K gene sets. All modules are discoverable via an advanced search engine that can filter by genes, metadata, and enrichment results. Analyses can also be browsed with an interactive workflow visualization tool, and users can communicate within OMICON using @mention functionality to support communal research on human brain gene coexpression networks.

17
In vivo multimodal lineage tracing of mammalian development by DeepTrack barcoding

Guo, C.; Jiang, J.; Wang, X.; Huang, X.; Zhang, S.; Shao, C.; Zhang, M.; Hu, X.; Yang, W.; Shang, F.; Wang, X.; Zhai, H.; Du, Q.; Liu, F.; He, D.; Liu, X.; Peng, G.; Cheng, S.; Zhang, Y.; Pei, D.; Pei, W.

2026-08-31 developmental biology 10.64898/2026.08.29.748052 medRxiv
Top 2%
6.0%
Show abstract

A comprehensive recording of cell fate transitions and underlying molecular changes remains a fundamental goal in developmental biology. Here, we present DeepTrack, a lineage tracing mouse model that integrates in situ cellular barcoding with high-throughput, single-cell multi-omics to simultaneously profile clonal fates, transcriptomic states, and chromatin accessibility. Using DeepTrack, we profiled clonal behaviors during gastrulation and early organogenesis, uncovered early fate priming within epiblast clones, and revealed clonal architecture within distinct regions of the nervous system. Embryo-wide multi-omic lineage tracing at single-cell resolution revealed transcriptional and epigenetic programs underlying fate commitment in neuromesodermal progenitors (NMPs). Clonal tracing with multi-omic profiles enabled inference of fate-associated gene-regulatory networks and identified the transcription factor Cdx2 as a key regulator of mesodermal specification in NMPs. Genetic perturbation of Cdx2 in chimeric embryos impaired paraxial mesoderm differentiation. Together, DeepTrack provides a versatile framework for decoding multimodal regulation of cell fate across diverse developmental contexts.

18
FlexiTAC enables controllable PROTAC linker generation across diverse structural settings using a Bayesian flow network with posterior guidance

Li, Y.; Zhao, Y.; Zhou, L.; Huang, C.; Xu, Q.; Chen, Y.; Qin, Z.; Fan, K.; Yang, J.; Cao, D.

2026-08-30 bioinformatics 10.64898/2026.08.26.747172 medRxiv
Top 2%
5.5%
Show abstract

Linker chemistry and conformation are central determinants of PROTAC activity, shaping ternary-complex geometry, cooperativity, target-lysine presentation and cellular permeability. Existing linker generators often lack explicit control over linker flexibility, require predefined attachment sites and linker lengths, or produce structures that demand substantial geometric correction, limiting their utility in practical PROTAC design. Here we introduce FlexiTAC, a Bayesian flow network that jointly generates linker atom types and coordinates from the warhead and E3-ligase-ligand contexts. We also assemble PROTAC-3D, a quality-controlled collection of 63,554 component-resolved PROTAC structures for model training, and PROTAC-Bench, which covers molecular quality, fragment preservation, geometric fidelity, conformational stability, fragment awareness, rediscovery and sampling efficiency. Compared to the best 3D baseline models, FlexiTAC improves validity by 12.0-12.7% and achieves the highest PoseBusters pass rate of 79.5%-80.0%. A differentiable guidance module shifted generated linkers along a conformational ensemble-derived rigidity axis without retraining the generator. In silico case studies further show that the model can accept crystal-derived, redocked or predicted structural inputs. Together, FlexiTAC, PROTAC-3D and PROTAC-Bench establish an integrated and reproducible framework for data-driven PROTAC linker design, combining controllable structure-conditioned generation with standardized training data and evaluation protocols. This framework expands the linker chemical and conformational space accessible to computational exploration, provides a foundation for future method development and enables the systematic generation of structure-conditioned linker designs with tunable conformational flexibility.

19
Paired-surface spatial mechanomics links tissue stiffness maps to spatial transcriptomics

Ong, H. T.; Lou, Y.; Turley, J.; Hengst, R. M.; Ramli, M. F. H.; Shen, X.; Marlena, J.; Zhu, J.; Li, R.; Chan, C. J.; Young, J. L.

2026-08-31 bioengineering 10.64898/2026.08.29.748050 medRxiv
Top 2%
5.5%
Show abstract

Tissue mechanics influence diverse biological processes, yet directly linking stiffness measurements to spatially resolved molecular states in intact tissues remains challenging. Here we developed a paired-surface spatial mechanomics approach to map Young's modulus by nanoindentation on a fresh tissue surface and co-register the stiffness grid with 10x Genomics Visium HD spatial transcriptome bins from the immediately adjacent, parallel surface. Applied to the mouse ovary, which has spatially distinct compartments and undergoes extracellular matrix remodeling with cycle and age, the workflow generated >2,900 matched measurements across 21 regions of interest. Nanoindentation at 50-m grid spacing enabled millimeter-scale stiffness maps while balancing acquisition time in fresh tissues, with ~92 4-m transcriptome bins assigned to each stiffness value. Global and compartment-specific analyses associated stiffer regions with lower elastic fiber programs and higher inflammatory signaling, with age-dependent differences. This correlative strategy integrates experimentally measured mechanics with spatial omics in fresh tissues.

20
Accurate and efficient prediction of protein conformations with ProtMonomer

Si, Y.; Zhang, S.; Chen, L.

2026-08-31 molecular biology 10.64898/2026.08.28.747824 medRxiv
Top 2%
5.4%
Show abstract

Deep learning-based protein structure prediction methods that leverage evolutionary information from multiple sequence alignments (MSAs), exemplified by AlphaFold2, have achieved remarkable accuracy. However, existing methods still struggle to predict challenging proteins, particularly those with novel folds or limited evolutionary information, and to recover alternative conformational states. Here we show that structure prediction models trained under different MSA-depth distributions corresponding to different levels of evolutionary information exhibit complementary generalization behaviors, and that a model trained on a mixture of these distributions can combine their complementary generalization strengths. Building on this insight, we developed ProtMonomer, a deep learning framework trained on MSA-depth distributions representing a broad range of evolutionary information levels to improve structure prediction. Across benchmarks comprising CASP15 targets, non-redundant experimentally determined structures, orphan proteins, and short peptides, ProtMonomer performed comparably to or better than leading methods, including AlphaFold2 and AlphaFold3, with particularly strong performance on challenging targets. For fold-switching proteins, ProtMonomer also recovered alternative conformational states more accurately than AlphaFold2 and AlphaFold3 across diverse homologous sequence sampling strategies. In addition to improving predictive accuracy, ProtMonomer substantially reduced inference cost through an efficient architecture, enabling high-throughput applications. Together, these findings provide insights into the generalization of evolution-informed structure prediction models and support ProtMonomer as an accurate and efficient framework for protein structure prediction.